Papers with English-centric corpora
Accelerating Multilingual Language Model for Excessively Tokenized Languages (2024.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have shown a significant degree of multilingual proficiency on a variety of tasks in multiple languages. |
| Approach: | They propose a framework to fine-tune a language model head and fine-track it while preserving its performance. |
| Outcome: | The proposed framework increases the generation speed by 1.7 while maintaining the performance of pre-trained multilingual models on target monolingual tasks. |
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs (2025.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) exhibit strong multilingual performance despite training on English-centric corpora. |
| Approach: | They propose to use Romanization as a potential bridge in multilingual processing . they propose to encode semantic concepts similarly across native and Romanized scripts . |
| Outcome: | The proposed model encodes semantic concepts across native and Romanized scripts, suggesting a shared underlying representation. |